<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="http://egonw.github.io/blog/feed/by_tag/json.xml" rel="self" type="application/atom+xml" /><link href="http://egonw.github.io/blog/" rel="alternate" type="text/html" /><updated>2026-09-29T06:27:22+00:00</updated><id>http://egonw.github.io/blog/feed/by_tag/json.xml</id><title type="html">chem-bla-ics</title><subtitle>Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.</subtitle><author><name>Egon Willighagen</name></author><entry><title type="html">Pulling out data as JSON from XHTML+RDFa</title><link href="http://egonw.github.io/blog/2010/09/10/pulling-out-data-as-json-from-xhtmlrdfa.html" rel="alternate" type="text/html" title="Pulling out data as JSON from XHTML+RDFa" /><published>2010-09-10T00:00:00+00:00</published><updated>2010-09-10T00:00:00+00:00</updated><id>http://egonw.github.io/blog/2010/09/10/pulling-out-data-as-json-from-xhtmlrdfa</id><content type="html" xml:base="http://egonw.github.io/blog/2010/09/10/pulling-out-data-as-json-from-xhtmlrdfa.html"><![CDATA[<p>I am keen on <a href="/blog/tag/rdfa">RDFa</a> and <a href="http://sv.wikipedia.org/wiki/Resource_Description_Framework">RDF <i class="fa-solid fa-recycle fa-xs"></i></a> in general; that should not be a surprise.
RDFa is a serialization of RDF triples embedded in (X)HTML. I recently posted about <a href="http://chem-bla-ics.blogspot.com/2010/08/xhtmlrdfa-chemical-examples.html">chemical examples of XHTML+RDFa</a>.
Now, the reason for putting data in HTML as RDFa is that we can easily pull it out again, e.g. with <a href="http://www.w3.org/2007/08/pyRdfa/">this distiller</a>.
But the fun goes on, and we can actually also run SPARQL directly on it, for example with RDFaDev which I
<a href="http://chem-bla-ics.blogspot.com/2010/07/scripts-logs-as-htmlrdfa-mix-free-text.html">recently blogged about</a>.</p>

<p>Now, consider we have all these nice visualization tools written in JavaScript which can visualize data from <a href="http://www.json.org/">JSON</a> sources,
the mashup requires a JSON serialization of that data embedded in HTML pages. Now, I have no experience with the cool JavaScript tools, and hope
someone can help me out here, but the JSON bit I already <a href="http://www.semanticoverflow.com/questions/587/is-there-a-web-service-that-allow-me-to-run-sparql-against-a-xhtmlrdfa-website">got help with before on SemanticOverflow</a>
(thanx to <a href="http://www.semanticoverflow.com/users/148/comment-bot">Comment Bot</a>!). The service mentioned no longer works, but there are plenty of alternatives.</p>

<p>Now, Peter is creating this nice data set about <a href="http://wwmm.ch.cam.ac.uk/blogs/murrayrust/?p=2596">green solvents from patents</a>, and it would be great of that data ends up online as RDFa, so that we can easily visualize the trends in solvent use over the years. But as I do not have this data as XHTML+RDFa yet, you will have to do with another example: boiling points.</p>

<p>So, let’s consider the data on <a href="https://egonw.github.io/cheminformatics.classics/classic1.html">this page <i class="fa-solid fa-recycle fa-xs"></i></a>, relating paraffin molecules to boiling points, and we’ll take a complexity descriptor (<em>w0</em>, Wiener descriptor) and the boilingpoint (<em>t0</em>). so we get this SPARQL query:</p>

<pre>PREFIX cc: &lt;http://github.com/egonw/cheminformatics.classics/1/#&gt;

SELECT * {
    ?mol cc:w0 ?w ;
         cc:p0 ?p .
}
</pre>

<p>Now, we want to run this query on the aforementioned page, so we add a FROM clause:</p>

<pre>PREFIX cc: &lt;http://github.com/egonw/cheminformatics.classics/1/#&gt;

SELECT *
FROM &lt;http://www.w3.org/2007/08/pyRdfa/extract?uri=http%3A%2F%2Fegonw.github.com%2Fcheminformatics.classics%2Fclassic1.html&amp;format=pretty-xml&amp;warnings=false&amp;parser=lax&amp;space-preserve=true&gt;
{
    ?mol cc:w0 ?w ;
         cc:p0 ?p .
}
</pre>

<p>Notice the use of the distiller here. This way, with a service like <a href="http://sparql.org/sparql.html">that on sparql.org</a>, we can get JSON returned. The result is a bit verbose, but that can perhaps be tuned:</p>

<pre>{
  "head": {
    "vars": [ "w" , "p" ]
  } ,
  "results": {
    "bindings": [
      {
        "w": { "datatype": "http://www.w3.org/2001/XMLSchema#integer" , "type": "typed-literal" , "value": "56" } ,
        "p": { "datatype": "http://www.w3.org/2001/XMLSchema#integer" , "type": "typed-literal" , "value": "4" }
      } ,
      {
        "w": { "datatype": "http://www.w3.org/2001/XMLSchema#integer" , "type": "typed-literal" , "value": "35" } ,
        "p": { "datatype": "http://www.w3.org/2001/XMLSchema#integer" , "type": "typed-literal" , "value": "3" }
      }
    ]
  }
}
</pre>

<p>The point is, I am sure at least one of my readers knows how to visualize the data in <a href="http://sparql.org/sparql?query=PREFIX+cc%3A+%3Chttp%3A%2F%2Fgithub.com%2Fegonw%2Fcheminformatics.classics%2F1%2F%23%3E%0D%0A%0D%0ASELECT+%3Fw+%3Fp%0D%0AFROM+%3Chttp%3A%2F%2Fwww.w3.org%2F2007%2F08%2FpyRdfa%2Fextract%3Furi%3Dhttp%253A%252F%252Fegonw.github.com%252Fcheminformatics.classics%252Fclassic1.html%26format%3Dpretty-xml%26warnings%3Dfalse%26parser%3Dlax%26space-preserve%3Dtrue%3E%0D%0A%7B%0D%0A++++%3Fmol+cc%3Aw0+%3Fw+%3B%0D%0A+++++++++cc%3Ap0+%3Fp+.%0D%0A%7D&amp;default-graph-uri=&amp;stylesheet=%2Fxml-to-html.xsl&amp;output=json&amp;force-accept=text%2Fplain">this JSON</a> with, for example, <a href="http://code.google.com/apis/chart/">Google Chart</a>, particularly, because all the mashing up is embedded in the just linked-to, though obscure, URL. And, if it helps, you can otherwise use the <a href="http://sparql.org/sparql?query=PREFIX+cc%3A+%3Chttp%3A%2F%2Fgithub.com%2Fegonw%2Fcheminformatics.classics%2F1%2F%23%3E%0D%0A%0D%0ASELECT+%3Fw+%3Fp%0D%0AFROM+%3Chttp%3A%2F%2Fwww.w3.org%2F2007%2F08%2FpyRdfa%2Fextract%3Furi%3Dhttp%253A%252F%252Fegonw.github.com%252Fcheminformatics.classics%252Fclassic1.html%26format%3Dpretty-xml%26warnings%3Dfalse%26parser%3Dlax%26space-preserve%3Dtrue%3E%0D%0A%7B%0D%0A++++%3Fmol+cc%3Aw0+%3Fw+%3B%0D%0A+++++++++cc%3Ap0+%3Fp+.%0D%0A%7D&amp;default-graph-uri=&amp;stylesheet=%2Fxml-to-html.xsl&amp;output=csv&amp;force-accept=text%2Fplain">CSV</a> or <a href="http://sparql.org/sparql?query=PREFIX+cc%3A+%3Chttp%3A%2F%2Fgithub.com%2Fegonw%2Fcheminformatics.classics%2F1%2F%23%3E%0D%0A%0D%0ASELECT+%3Fw+%3Fp%0D%0AFROM+%3Chttp%3A%2F%2Fwww.w3.org%2F2007%2F08%2FpyRdfa%2Fextract%3Furi%3Dhttp%253A%252F%252Fegonw.github.com%252Fcheminformatics.classics%252Fclassic1.html%26format%3Dpretty-xml%26warnings%3Dfalse%26parser%3Dlax%26space-preserve%3Dtrue%3E%0D%0A%7B%0D%0A++++%3Fmol+cc%3Aw0+%3Fw+%3B%0D%0A+++++++++cc%3Ap0+%3Fp+.%0D%0A%7D&amp;default-graph-uri=&amp;stylesheet=%2Fxml-to-html.xsl&amp;output=tsv&amp;force-accept=text%2Fplain">TSV</a> output. The output of that is even more simple (CSV):</p>

<pre>w,p
56,4
286,9
35,3
220,8
20,2
84,5
10,1
165,7
120,6
</pre>

<p>The first one who can use one of the above URLs to extract the data from that XHTML+RDFa page to create a scatter plot in a HTML page with some JavaScript library, wins a free mention in my blog! ;)</p>]]></content><author><name>Egon Willighagen</name></author><category term="chemistry" /><category term="html" /><category term="rdfa" /><category term="sparql" /><category term="json" /><summary type="html"><![CDATA[I am keen on RDFa and RDF in general; that should not be a surprise. RDFa is a serialization of RDF triples embedded in (X)HTML. I recently posted about chemical examples of XHTML+RDFa. Now, the reason for putting data in HTML as RDFa is that we can easily pull it out again, e.g. with this distiller. But the fun goes on, and we can actually also run SPARQL directly on it, for example with RDFaDev which I recently blogged about.]]></summary></entry></feed>